Papers with mechanistic interpretability tools

3 papers
From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent research in mechanistic interpretability has revealed that Large Language models contain disentangled, human-understandable components.
Approach: They propose a framework that first identifies causal task features through frequency recall and interventional filtering, then selects “Feature-Resonant Data” that maximally activates task features for fine-tuning.
Outcome: The proposed framework outperforms existing models on mathematical reasoning, summarization, and translation tasks while using only 50% of the data.
Preference Tuning For Toxicity Mitigation Generalizes Across Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Detoxifying multilingual Large Language Models (LLMs) has become crucial due to their increasing global use.
Approach: They propose to use English preference tuning to study cross-lingual detoxification of LLMs.
Outcome: The proposed method reduces toxicity in multilingual LLMs by reducing the probability of mGPT-1.3B generating toxic continuations across 17 languages.
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models excel at extracting and reasoning about information from images, yet their capacity to leverage internal knowledge about specific entities remains underexplored.
Approach: They propose a dataset which allows separating entity recognition and question answering . they hypothesize that this decline arises from limitations in how information flows from image tokens to query tokens.
Outcome: The proposed model performance drops when the entity is presented visually rather than textually.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations